Lightweight, highly optimized CPU runtime for GGUF models and embeddings.
DetailsPrivate, local AI desktop app — run open LLMs fully offline with a coding agent, knowledge base and voice.
Detailshtop for your local AI
DetailsModel swapping for llama.cpp
Details